Papers with extended dataset

5 papers
Transfer Learning from Transformers to Fake News Challenge Stance Detection (FNC-1) Task (2020.lrec-1)

Copied to clipboard

Challenge: In the last two years, significant improvements have occurred in NLP with the development of large language models using contextualized word embeddings based on the Google Transformer architecture.
Approach: They performed experiments on data from the Fake News Challenge stage 1 (FNC-1) they used BERT sentence embeddings as a model feature and BERT, XLNet, and RoBERTa transformers to fine-tune them.
Outcome: The proposed model outperforms the winner's system on class-wise F1 scores and achieves state-of-the-art on the stance detection task.
Towards Cross-Lingual Explanation of Artwork in Large-scale Vision Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: LVLMs are increasingly capable of responding in multiple languages . however, there is a lack of evaluation tools for LVLs that handle multiple languages.
Approach: They used an extended dataset in multiple languages to evaluate LVLMs' ability to generate explanations in multiple language combinations.
Outcome: The proposed dataset in multiple languages evaluates LVLMs' ability to generate explanations in other languages.
Embedding Hallucination for Few-shot Language Fine-tuning (2022.naacl-main)

Copied to clipboard

Challenge: Existing methods for fine-tuning pre-trained language models can cause severe over-fitting.
Approach: They propose an Embedding Hallucination method which generates auxiliary embedding-label pairs to expand the fine-tuning dataset.
Outcome: The proposed method outperforms current fine-tuning methods in a wide range of language tasks.
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models (2026.acl-long)

Copied to clipboard

Challenge: a recent study evaluated large audio-language models against jailbreak attacks . a new benchmark is being developed to evaluate LAM safety against jailbreaking attacks based on temporal and semantic nature of speech .
Approach: They propose a benchmark to evaluate LAM jailbreak vulnerabilities in adversarial audio prompts . they use a dataset of 1,495 adversarials to evaluate their performance .
Outcome: The proposed benchmark evaluates state-of-the-art LAMs against jailbreak attacks . it demonstrates that even small, semantically preserved perturbations can reduce safety .
Bilingual Zero-Shot Stance Detection (2025.acl-long)

Copied to clipboard

Challenge: a study focuses on noun-phrase and claim targets within bilingual ZSSD scenarios . a dataset focusing on claim targets with a low occurrence of shared words is also explored .
Approach: They use a bilingual bilingual ZSSD dataset to investigate the use of zero-shot stance detection.
Outcome: The proposed dataset is the first to examine this difficult setting in bilingual ZSSD . it focuses on noun-phrase and claim targets within in-domain and out-of-domain bilingual scenarios .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations